Papers with computer vision
Copied to clipboard
| Challenge: | EMNLP 2025 Industry Track highlights key insights, novel research trends and challenges encountered in practical language technology applications. |
| Approach: | Kai Chen will present the technical advances behind the open-source Intern-series large models . he will highlight how models acquire expert-level skills in specialized domains . |
| Outcome: | This talk will highlight the technical advances behind the open-source Intern-series models . it will highlight how models acquire expert-level skills in specialized domains while retaining broad generalization ability. |
Copied to clipboard
| Challenge: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Approach: | Silver es provides high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Outcome: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
Copied to clipboard
| Challenge: | In-store users only need to take a picture or scan the barcode of the product of interest, and then the user can talk to the assistant about the product. |
| Approach: | They present a mobile-based intelligent shopping assistant that is designed to improve shopping experience in physical stores. |
| Outcome: | The proposed system can improve shopping experience in physical stores by leveraging advanced techniques in computer vision, speech processing, and natural language processing. |
Copied to clipboard
| Challenge: | Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events. |
| Approach: | They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations. |
| Outcome: | The proposed methods are based on the models of large language models (LLMs) and their opaque nature. |
Copied to clipboard
| Challenge: | supervised machine learning is based on learning in isolation, a single predictive model for a task using a dataset. |
| Approach: | They present an overview of modern transfer learning methods in natural language processing . they review examples and case studies on how models can be integrated and adapted . |
| Outcome: | The proposed methods improve upon the state-of-the-art on a wide range of NLP tasks. |
Copied to clipboard
| Challenge: | Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored. |
| Approach: | They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
| Outcome: | The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
Copied to clipboard
| Challenge: | This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions . |
| Approach: | This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments . |
| Outcome: | This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough . |
Copied to clipboard
| Challenge: | Recent advances in machine learning have led to the use of contrastive loss for representation learning. |
| Approach: | They propose to use batch-softmax contrastive loss to train pairwise sentence embeddings . they propose to take a batch-softermax contrastitive loss and train it with different loss functions . |
| Outcome: | The proposed model improves on a number of datasets and pairwise sentence scoring tasks. |
Copied to clipboard
| Challenge: | Existing studies on image caption generation in English focus on Western languages, ignoring Semitic and Middle-Eastern languages like Arabic, Hebrew, Urdu and Persian. |
| Approach: | They propose to leverage the critical dependency of Arabic to generate Arabic captions using root-word based Recurrent Neural Network and Deep Neural networks. |
| Outcome: | The proposed model outperforms English-Arabic translated captions on a dataset from newspapers in the Middle East. |
Copied to clipboard
| Challenge: | Existing adversarial attacks can cause LLMs to make wrong predictions on downstream tasks or generate harmful content misaligned with human values. |
| Approach: | They propose to use randomized smoothing to add noise to the input and then make predictions based on these denoised versions. |
| Outcome: | The proposed method surpasses existing methods in both empirical and certified robustness in defending against adversarial perturbations for both downstream tasks and human alignments (i.e., jailbreak attacks). |
Copied to clipboard
| Challenge: | Experimental evaluations demonstrate FID score of 8.03 on the COCO-30K dataset, marking our model as the top open source performer in terms of measurable image generation quality. |
| Approach: | They propose a latent diffusion-based model that combines image prior and latent diffusive techniques to create a text-to-image architecture. |
| Outcome: | The proposed model achieves the highest FID score among open-source models . it is compared with the state-of-the-art models on the COCO-30K dataset . |
Copied to clipboard
| Challenge: | Obtaining high-quality labeled data that accurately represents complexity of real-world scenarios can be expensive, time-consuming, or even impractical. |
| Approach: | They propose to use Fréchet Inception Distance to measure distance between judged items and retrieved results. |
| Outcome: | The proposed method improves on a MS MARCO dataset and TREC Deep Learning Tracks query sets. |
Copied to clipboard
| Challenge: | In contrast, adversarial training has been used in computer vision to improve models’ robustness due to the discrete nature of text. |
| Approach: | They propose a way to generate adversarial samples by using pseudo-labeled in-domain text data to train a seq2seq model for adversarials and combine it with paraphrase detection. |
| Outcome: | The proposed model generates realistic and relevant adversarial samples compared to other state-of-the-art models and recovers up to 70% of errors. |
Copied to clipboard
| Challenge: | Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content. |
| Approach: | They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future . |
| Outcome: | The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future . |
Copied to clipboard
| Challenge: | Transformer model has been a de-facto standard in natural language processing, but it is limited to images, text, and/or sequence data. |
| Approach: | They propose to use a multimodal large language model architecture to handle biomedical graphs such as protein structure and chemical molecules to improve its performance. |
| Outcome: | The proposed architecture can handle multiple data types for biomedical graphs such as protein structure and chemical molecules. |
Copied to clipboard
| Challenge: | Existing approaches to computer vision require task-specific modifications and training from scratch. |
| Approach: | They propose a method that can be applied to any task in NLP and propose to open-source it. |
| Outcome: | The proposed method outperforms the state-of-the-art on six text classification tasks, reducing error by 18-24% on majority of datasets. |
Copied to clipboard
| Challenge: | Large pre-trained Vision-Language Models (VLMs) have revolutionized downstream vision-language tasks including classification, object detection, and segmentation. |
| Approach: | They propose to search for text prompts at the word level rather than optimizing continuous textual embeddings to boost adversarial robustness. |
| Outcome: | Experiments show that the proposed method outperforms hand-engineered prompts with average gains of +4.9% and +5.8%. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities. |
| Approach: | They propose an algorithm which learns a joint Concept-Vision-Language embedding for questions which require common sense knowledge from external structured content. |
| Outcome: | The proposed model is based on the Outer Knowledge-VQA and VQA datasets. |
Copied to clipboard
| Challenge: | Existing tools for working with scientific documents are limited and documents are often in difficult-to-use PDF formats. |
| Approach: | They propose an open-source Python toolkit for analyzing and processing visually-rich scientific documents. |
| Outcome: | PaperMage provides turn-key recipes for common scientific document processing use-cases. |
Copied to clipboard
| Challenge: | Existing methods to overcome overfitting in text learning do not consider dimensionality . dimensionalization is important for deep neural networks to overcome the problem . |
| Approach: | They propose a saliency map-based approach to overcome overfitting in text learning . they propose augmentation regularization methods such as Dropout and Mixup to improve regularization . |
| Outcome: | Empirical results show that the proposed approach overcomes overfitting in text learning . dropout and mixup methods are effective in enhancing regularization . |
Copied to clipboard
| Challenge: | Large-scale pretraining and task-specific fine-tuning are now the standard methodology for many tasks in computer vision and natural language processing. |
| Approach: | They propose to combine two types of vision and language BERTs to create a theoretical framework that can be unified under different theoretical frameworks. |
| Outcome: | The proposed models can be classified into single-stream or dual-stream encoders and are unified under a single theoretical framework. |
Copied to clipboard
| Challenge: | Recent advances in pre-trained language models have been limited when fine-tuned on small datasets. |
| Approach: | They propose to add contrastive learning to prompt-based fine-tuning to improve model performance. |
| Outcome: | The proposed approach outperforms other methods on multiple text classification benchmarks. |
Copied to clipboard
| Challenge: | Quantization-aware training (QAT) is a promising method to lower the implementation cost and energy consumption. |
| Approach: | They propose a method for fast converging QAT of pre-trained Transformers using a layer-wise signal propagation method with the intact signal from the teacher. |
| Outcome: | The proposed method achieves superior accuracy with significantly lower fine-tuning iterations on well-known Transformers of natural language processing as well as computer vision compared to the state-of-the-art methods. |
Copied to clipboard
| Challenge: | Previously, domain adaptation approaches to bilingual tasks were proposed . we show that simple adaptation process involving only unlabeled text is highly effective . |
| Approach: | They propose a method for domain adaptation of bilingual word embeddings using unlabeled data . they then tailor a semi-supervised classification method from computer vision to these tasks . |
| Outcome: | The proposed method improves on two bilingual tasks using unlabeled data. |
Copied to clipboard
| Challenge: | Recent studies have shown that Mamba models can be used for multiple domains, including NLP, long-range sequence processing, and computer vision. |
| Approach: | They add a third view and show that Mamba models can be viewed as attention-driven models. |
| Outcome: | The proposed model can be viewed as attention-driven and empirically compare it to the attention-based models of transformers. |
Copied to clipboard
| Challenge: | Recent work on few-shot learning addresses the problem of learning based on a small amount of training data. |
| Approach: | They adapt the Amazon Review Sentiment Classification (ARSC) text dataset for few-shot learning . they train a single binary classifier to learn all few- shot classes jointly . |
| Outcome: | The proposed approach outperforms most published results on the ARSC text dataset . the results suggest that the classes in the AR SC few-shot task are very similar to each other . |
Copied to clipboard
| Challenge: | Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud. |
| Approach: | They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing. |
| Outcome: | The proposed models perform better on dialog act classification tasks while maintaining high accuracy. |
Copied to clipboard
| Challenge: | a novel approach to contrastive learning for language understanding is not fully explored . contrastive training has been widely applied to self-supervised representation learning . |
| Approach: | They propose a label anchored contrastive learning approach for language understanding using a class label. |
| Outcome: | The proposed approach improves on GLUE and CLUE benchmarks by 4.1% compared to the state-of-the-art approaches . the proposed approach also improves under the few-shot and data imbalance settings . |
Copied to clipboard
| Challenge: | Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud. |
| Approach: | They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing. |
| Outcome: | The proposed models perform better on dialog act classification tasks while maintaining high accuracy. |
Copied to clipboard
| Challenge: | a mixed methods approach is feasible for the identification of scribes and authors in handwritten documents. |
| Approach: | They propose a mixed methods approach to the identification of scribes and authors in handwritten documents . they use a software tool which combines linguistic insights and computer vision techniques . |
| Outcome: | The proposed tool can be used to identify scribes and authors in handwritten documents. |
Copied to clipboard
| Challenge: | In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem. |
| Approach: | They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining. |
| Outcome: | The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes. |
Copied to clipboard
| Challenge: | Existing low-resource datasets that challenge neural networks cause over-estimated performance, despite promising yet saturated results in high-res settings. |
| Approach: | They propose a benchmark Achilles-Bench to better evaluate the learning ability of neural networks in low-resource settings. |
| Outcome: | The proposed benchmarks show that even pre-trained language models show performance drops on NLP tasks. |
Copied to clipboard
| Challenge: | To explain NLP models, importance measures are often used to inform input tokens are important for making a prediction. |
| Approach: | They propose a faithfulness metric that masks allegedly important tokens and retrains the model. |
| Outcome: | The proposed metric is based on LSTM-attention models and RoBERTa models. |
Copied to clipboard
| Challenge: | Neural networks are notoriously data-hungry, resulting in ungrammatical texts . data augmentation requires a specific design for a structurally rich input format . |
| Approach: | They propose to selectively augment a training set with new data by adding and varying two specific lexical categories, i.e. proper and common nouns. |
| Outcome: | The proposed approach selectively augments a training set with new data by adding and varying two specific lexical categories, i.e. proper and common nouns. |
Copied to clipboard
| Challenge: | Existing work on integrating graph problems into generative language modeling framework remains limited. |
| Approach: | They propose an LLM with instructions based on natural language to perform graph tasks. |
| Outcome: | The proposed model surpasses all GNN baselines on ogbn-arxiv, Cora and PubMed datasets and sheds light on generative LLMs as new foundation model for graph machine learning. |
Copied to clipboard
| Challenge: | Recent studies have used metric-based learning in computer vision but not slot tagging. |
| Approach: | They propose a metric-based learning architecture that extends relation networks by leveraging pretrained contextual embeddings such as ELMO and BERT and by using attention mechanism. |
| Outcome: | The proposed method outperforms state-of-the-art methods on SNIPS data on a slot tagging task with a large amount of hand-labeled data. |
Copied to clipboard
| Challenge: | Pre-training large language models can be expensive and wasteful. |
| Approach: | They propose a method which can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and a two-stage learning method to further accelerate the pre-training. |
| Outcome: | The proposed method can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and significantly improve the pre-training efficiency of the large model. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines. |
| Approach: | They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data. |
| Outcome: | The proposed model can improve when questions are modified to include more details. |
Copied to clipboard
| Challenge: | a growing popularity of deep-learning models makes model understanding more important . feature attribution methods have shown promising results in computer vision but are not trivial . |
| Approach: | They propose a gradient-based feature attribution method that smooths gradients by aggregating similar reference texts derived from language model embeddings. |
| Outcome: | The proposed method outperforms existing methods on public datasets and key words detection tasks. |
Copied to clipboard
| Challenge: | a new method for analyzing prosody in sign languages uses the velocity profile of the hands . the velocity profiles of hand movements can be used to analyse prosodic structure . |
| Approach: | They propose a method for extracting velocity information from unlabeled video of sign language using CoTracker. |
| Outcome: | The proposed method can extract prosodic information from unlabeled video clips. |
Copied to clipboard
| Challenge: | a new schema for NLP knowledge about tasks, datasets and metrics is proposed. |
| Approach: | They propose a new schema that represents knowledge about tasks, datasets and metrics in the NLP domain. |
| Outcome: | The proposed framework can be automatically built into scientific leaderboards . the proposed system achieves reasonable results for all relation types on this small-scale graph . |
Copied to clipboard
| Challenge: | Story visualization is an underexplored task that requires a generative model to generate images . prior work has focused on image generation but there is room for improvement . |
| Approach: | They propose to add a dual learning framework to reinforce semantic alignment between story and generated images and a copy-transform mechanism to model sequentially-consistent story visualization. |
| Outcome: | The proposed models outperform text-to-image synthesis models on the story visualization task . the proposed models also improve visual quality, coherence and relevance . |
Copied to clipboard
| Challenge: | a long-term goal of artificial intelligence is to have an agent execute commands through natural language. |
| Approach: | They propose to use a dataset to compare commands written in natural language for self-driving cars with other datasets. |
| Outcome: | The proposed task is a challenging one and shows promising results, the authors argue . the talk2car dataset compares with similar datasets and shows that the proposed task requires additional research in natural language processing and computer vision. |
Copied to clipboard
| Challenge: | Sentence embedding models are typically trained using contrastive learning (CL) using human annotations directly or by repurposing other annotated datasets. |
| Approach: | They propose to use generative language models to generate CL data using annotated data. |
| Outcome: | The proposed method outperforms the previous best unsupervised method by 1.8 points and SimCSE, a strong supervised baseline by 0.3 points on the semantic text similarity (STS) benchmark. |
Copied to clipboard
| Challenge: | a recent study evaluated the impact of differential privacy on fairness across four tasks. |
| Approach: | They evaluate the impact of differential privacy on fairness across four diverse tasks . they train (,)-differentially private models with empirical risk minimization . |
| Outcome: | The proposed model shows that differential privacy increases performance differences between groups . the model also reduces performance differences in the robust setting . |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of the Retentive Networks have not yet been conducted. |
| Approach: | They propose a retention mechanism that integrates the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models. |
| Outcome: | The proposed retention mechanism combines the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models. |
Copied to clipboard
| Challenge: | Existing methods for enhancing in-context emotion classification fail to include spatial relationships between different people and facial features within a single face. |
| Approach: | They propose a set-of-vision prompting approach that uses spatial information to mark targets precisely. |
| Outcome: | The proposed approach improves face count and emotion categorization while preserving the enriched image context. |
Copied to clipboard
| Challenge: | Graph Neural Networks (GNNs) are used to train neural networks to detect fake news based on context-based methods. |
| Approach: | They propose to combine the two by applying pre-training of Graph Neural Networks (GNNs) in the domain of context-based fake news detection. |
| Outcome: | The proposed methods show that transfer learning does not lead to significant improvements over training a model from scratch in the domain of context-based fake news detection. |
Copied to clipboard
| Challenge: | Existing studies of language-in-interaction focus on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. |
| Approach: | They propose to use CHICA to analyze child-caregiver conversations at home . they use mobile, lightweight eye-tracking and head motion detection to optimize the naturalness of the recordings. |
| Outcome: | The proposed corpus of child-caregiver conversations at home was compared with a previous corpus based on a set of conversations between children aged 7, 9, and 11 years old. |
Copied to clipboard
| Challenge: | SupCL-Seq extends contrastive learning from computer vision to sequence classification tasks. |
| Approach: | They propose a supervised alternative to Masked Language Modeling (MLM) that extends contrastive learning to sequence optimization in NLP by altering the dropout mask probability in standard Transformer architectures. |
| Outcome: | The proposed method leads to large gains on the GLUE benchmark, including 6% absolute improvement on CoLA, 5.4% on MRPC, 4.7% on RTE and 2.6% on STS-B. |
Copied to clipboard
| Challenge: | Eye4Ref is a rich multimodal dataset of eye-movement recordings from referentially complex situated settings. |
| Approach: | They present a rich multimodal dataset of eye-movement recordings from situated settings . they use linguistic labels, saccadic movement parameters and symbolic knowledge representations . |
| Outcome: | The Eye4Ref dataset is an annotated multimodal dataset from three eyetracking studies on reference resolution and disambiguation tasks in situated settings. |
Copied to clipboard
| Challenge: | Language-based learning models (LLMs) support long context-lengths but their effectiveness in handling long-term information gradually declines with input length. |
| Approach: | They propose a Language Repository (LangRepo) that maintains concise and structured information as an interpretable representation. |
| Outcome: | The proposed framework is evaluated on zero-shot visual question-answering benchmarks. |
Copied to clipboard
| Challenge: | Technical logbooks are a challenging and under-explored text type in automated event identification. |
| Approach: | They propose a feedback strategy that resamples the training data based on its error in the prediction process. |
| Outcome: | The proposed approach provides the best results for four different neural network models trained across a suite of technical logbook datasets from distinct technical domains. |
Copied to clipboard
| Challenge: | a new method for learning unsupervised sentence embeddings is proposed . unsup-SimCSE is biased because of the length information encoded into the sentence embeds . |
| Approach: | They propose a new unsupervised sentence embedding method that uses dropout to obtain positive pairs from a pre-trained Transformer encoder. |
| Outcome: | The proposed method outperforms the state-of-the-art unsup-SimCSE on a STS task. |
Copied to clipboard
| Challenge: | SlimFit reduces the memory requirements of transformer-based models by analyzing their training dynamics and freezing less-contributory layers during fine-tuning. |
| Approach: | They propose a tool that analyzes transformer-based models and freezes less-contributory layers during fine-tuning to reduce the overall on-device memory usage. |
| Outcome: | SlimFit reduces the memory requirements of transformer-based models by analyzing their training dynamics and freezing less-contributory layers during fine-tuning. |
Copied to clipboard
| Challenge: | Transferability estimation has been a topic of great interest in computer vision fields . a lack of a comprehensive comparison between these estimation methods is a problem . |
| Approach: | They conduct a thorough survey of existing methods to find the most suitable model . they also outline difficulties of consideration of training details and applicability to text generation . |
| Outcome: | The proposed methods perform well with superiorities in effectiveness and efficiency. |
Copied to clipboard
| Challenge: | a new study shows that general abusive language classifiers are reliable in detecting explicit abuse but fail to detect more subtle abuses. |
| Approach: | They propose an interpretability technique to quantify the sensitivity of a trained model to new data . they propose a degree of explicitness metric to suggest out-of-domain unlabeled examples . |
| Outcome: | The proposed interpretability technique is useful for predicting the generalizability of the model on new data. |
Copied to clipboard
| Challenge: | Recent studies have shown that AI is unfair in many real-world applications such as computer vision and recommendations. |
| Approach: | They propose to use a benchmark dataset to study the fairness of dialogue systems to understand their bias. |
| Outcome: | The proposed methods reduce the bias in dialogue systems significantly. |
Copied to clipboard
| Challenge: | Visual representation learning has been a cornerstone in computer vision for decades. |
| Approach: | They propose a visual representation tailored for visual reasoning that provides instance-level world knowledge and detailed attributes that are essential for visual reason. |
| Outcome: | The proposed visual tables outperform existing models on 11 visual reasoning benchmarks. |
Copied to clipboard
| Challenge: | Despite advances in computer vision, its application on language input still needs to be explored despite its feasibility. |
| Approach: | They propose a universal domain adaptation (uniDA) benchmark for natural language that offers thorough viewpoints of the model’s generalizability and robustness. |
| Outcome: | The proposed model can handle spoken language in the real world while also detecting unprocessable inputs from the target domain. |
Copied to clipboard
| Challenge: | Task-agnostic data augmentations have proven widely effective in computer vision, even on pretrained models. |
| Approach: | They examine the effects of two types of task-agnostic data augmentation on pretrained transformers using 5 classification tasks and 6 datasets. |
| Outcome: | The proposed techniques improve performance on 5 classification tasks, 6 datasets, and 3 variants of modern pretrained transformers. |
Copied to clipboard
| Challenge: | Existing Transformer Architecture Search methods are limited to computer vision and natural language processing tasks. |
| Approach: | They propose a Transformer Architecture Search proxy that measures trainability and expressivity of Transformer networks separately and integrates it into an effective regularized evolution framework to demonstrate its efficacy. |
| Outcome: | The proposed proxy can achieve higher correlation with the true performance of Transformer networks on computer vision and natural language processing tasks. |
Copied to clipboard
| Challenge: | Understanding images and text together is an important aspect of cognition and building advanced AI systems. |
| Approach: | They propose to derive joint inference about a given image-text modality and compile a question-answering corpus using an image and a reading passage. |
| Outcome: | The proposed method has better baseline performance but is still far behind human performance. |
Copied to clipboard
| Challenge: | Existing approaches to graph representation only consider the local neighbors, sacrificing the Transformer’s ability to attend to elements at any distance. |
| Approach: | They propose a dual-encoding Transformer architecture that uses a structural encoder and a semantic encoder to seek for semantically relevant nodes. |
| Outcome: | The proposed architecture achieves superior performance compared to state-of-the-art attention-based methods on complex relational graphs like KGs and citation networks. |
Copied to clipboard
| Challenge: | Text-to-image generation models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities. |
| Approach: | They propose a RusCode benchmark to evaluate the quality of text-to-image generation containing elements of the Russian cultural code. |
| Outcome: | The proposed model is based on 1250 text prompts in Russian and their translations into English. |
Copied to clipboard
| Challenge: | Existing studies focus on the text modality or are limited to specific tasks. |
| Approach: | They propose a framework to teach Large Vision-Language Models to selectively utilize retrieved information and improve their robustness against irrelevant or misleading references. |
| Outcome: | The proposed framework improves LVLMs’ ability to utilize retrieved multimodal references and their robustness against irrelevant or misleading information. |
Copied to clipboard
| Challenge: | Recent work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is given to the quality of vision models. |
| Approach: | They develop a selective attention model to study the patch-level contribution of an image in multimodal machine translation. |
| Outcome: | The proposed model is able to learn translation from the visual modality on probing tasks and is compared with existing models. |
Copied to clipboard
| Challenge: | Existing studies have tried to introduce discrete or Gaussian-based latent variables to address the one-to-many problem, but the diversity is limited. |
| Approach: | They propose a diffusion model to enhance the diversity of dialogue generation by using continuous latent variables instead of discrete ones. |
| Outcome: | The proposed model greatly enhances diversity of dialog response while keeping the coherence. |
Copied to clipboard
| Challenge: | Natural language processing (NLP) applications are growing rapidly due to discrete nature of texts. |
| Approach: | They propose to connect discrete perturbations with continuous perturbations to help understand discrete ones in NLP models. |
| Outcome: | The proposed method surpasses methods used in discrete perturbation measuring and can be generalized to different datasets, perturbation methods. |
Copied to clipboard
| Challenge: | Existing methods for adversarial samples are poorly applied in computer vision . however, textual adversarials are still vulnerable to small perturbations . |
| Approach: | They propose a framework to extend existing adversarial attack methods to textual adversarials by adding optimized perturbations to embedding layer and amplifying them in forward propagation process. |
| Outcome: | The proposed framework achieves better performance even using proxy gradient information and produces more fluent and grammatical adversarial samples compared to baseline methods. |
Copied to clipboard
| Challenge: | Existing adversarial defense methods for natural language processing still pose challenges to adversarials. |
| Approach: | They propose a novel adversarial defense method that incorporates a diffusion layer as a denoiser between the encoder and the classifier. |
| Outcome: | The proposed method improves over existing adversarial defense methods and achieves state-of-the-art performance against black-box and white-box adversarials. |
Copied to clipboard
| Challenge: | Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they use a single dataset-wise hyperparamter to control the strength of noise resistance. |
| Approach: | They propose to change single dataset-wise hyperparameters of noise resistance to be instance-wise. |
| Outcome: | The proposed frameworks increase noise-robustness on noisy and corrupted NLP datasets. |
Copied to clipboard
| Challenge: | Generative AI has demonstrated unprecedented creativity in the field of computer vision, yet such phenomena have not been observed in the realm of literary creation. |
| Approach: | They propose a framework for unleashing the creativity of large language models (LLMs) they assign LLMs to different roles involved in real-world scenario, they write . |
| Outcome: | The proposed framework outperforms baselines in terms of coherence, relevance, interestingness and overall quality on automatically generated screenplays. |
Copied to clipboard
| Challenge: | Transformer-based models are stretched to enormous sizes, requiring increasingly larger training datasets and unsustainable amount of compute resources. |
| Approach: | They propose an alternative compatibility function for the Transformer-based attention mechanism that exploits an overlap in the learned representation of the traditional scaled dot-product attention mechanism. |
| Outcome: | The proposed model achieves 79.36 on the GLUE benchmark against 78.74 for the traditional implementation and reduces the number of trainable parameters by 6%. |
Copied to clipboard
| Challenge: | Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts. |
| Approach: | They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels. |
| Outcome: | The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels. |
Copied to clipboard
| Challenge: | Depending on the size of transformer-based models, they can be restricted from deployment in resource-constrained environments. |
| Approach: | They propose to combine neural architecture search and network pruning techniques to generate and train weight-sharing super-networks that contain efficient transformer-based models. |
| Outcome: | The proposed model achieves high-performing, high-performance subnetworks on the general language understanding evaluation and the Stanford Question Answering Dataset. |
Copied to clipboard
| Challenge: | Prior work has proposed to augment Transformer model with the capability of skimming tokens to improve its computational efficiency. |
| Approach: | They propose to add a parameterized predictor before each layer that learns to make the skimming decision. |
| Outcome: | The proposed model achieves 10.97x speedup on GLUE benchmark compared with BERT-base baseline with less than 1% accuracy degradation. |
Copied to clipboard
| Challenge: | Extensive experiments demonstrate that treating attention as a feature map and applying convolution as . a processing method significantly enhances Transformer performance. |
| Approach: | They propose to use the convolution operator to mimic the processing methods in computer vision to treat attention as a feature map and apply it to neighboring attention scores across different heads. |
| Outcome: | The proposed model can be adapted to various attention-related models and achieves high performance. |
Copied to clipboard
| Challenge: | Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence. |
| Approach: | They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments. |
| Outcome: | This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work . |
Copied to clipboard
| Challenge: | Existing methods to detect adversarial text inputs are limited in performance and are not detectable via spell checkers. |
| Approach: | They propose a model-agnostic detector of adversarial text examples that detects patterns in the logits of the target classifier when perturbing the input text. |
| Outcome: | The proposed detector improves the state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks. |
Copied to clipboard
| Challenge: | Neural machine translation models with tens and even more than a hundred blocks have shown effectiveness in image recognition. |
| Approach: | They propose a two-stage approach with three specially designed components to construct deeper NMT models. |
| Outcome: | The proposed approach improves on WMT14 EnglishGerman and EnglishFrench translation tasks. |
Copied to clipboard
| Challenge: | Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions . |
| Approach: | They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances . |
| Outcome: | The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets . |
Copied to clipboard
| Challenge: | a crowdsourced method to evaluate saliency methods in NLP is proposed . saliencies are difficult for humans to understand, and can cause psychological harm . |
| Approach: | They propose a method to evaluate saliency methods in NLP by crowdsourcing . they recruited 800 crowd workers and empirically evaluated seven salience methods . |
| Outcome: | The proposed method evaluates saliency methods on two datasets using crowdsourced data . it shows that the results are comparable to existing methods on NLP and CV fields . |
Copied to clipboard
| Challenge: | Existing acceleration methods for text generation ignore the importance of the distribution of sampling steps, resulting in slow sampling rates. |
| Approach: | They propose a technique to accelerate diffusion models for text generation without additional training by using a Bayesian optimization approach. |
| Outcome: | The proposed technique achieves 400x acceleration even with minimal sampling steps after down to less than 1 minute of optimization yielding a competitive performance even with minimum sampling steps. |
Copied to clipboard
| Challenge: | Existing model-based channel prediction methods suffer from limited accuracy due to imperfect temporal modeling, while existing AI-based methods suffers from limited generalization due to inadequate training strategies. |
| Approach: | They propose a generative pre-trained language model for channel prediction based on channel correlation and train it based upon transformer decoder architecture. |
| Outcome: | The proposed model can learn various channel characteristics and perform impressive tasks across multiple dimensions. |
Copied to clipboard
| Challenge: | Textbooks lack visuals that support student learning, but many lack them . e-textbooks lack such visuals, and many lack these visuals . |
| Approach: | They propose to use vision-language models to automatically enhance textbooks with images from the web. |
| Outcome: | The proposed model improves textbooks with images from the web while allowing for better pedagogical value. |
Copied to clipboard
| Challenge: | emergence of large language models has significantly transformed the applications of deep learning methods in natural language processing. |
| Approach: | They propose to improve LLMs' generalization by optimizing entire models in parameter space by learning entire simplexes of continous prefixes. |
| Outcome: | The proposed method improves generalization of large language models in the scarce data regime. |
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |
Copied to clipboard
| Challenge: | In this work, we explore the extension of prototypical networks to natural language processing. |
| Approach: | They propose a weighted similarity measure that enhances the similarity computation by focusing on informative dimensions of pre-trained sentence embeddings. |
| Outcome: | The proposed method improves predictive performance on AG News and RT Polarity datasets and the rationale-based recurrent convolutions. |
Copied to clipboard
| Challenge: | Recent advances in the field of computer vision have enabled more effective and sophisticated interactions between humans and machines. |
| Approach: | They propose a reasoning-based object detection paradigm that leverages state-of-the-art multi-modal models and open-vocabulary object detectors to perform reasoning within the context of the user’s instructions and the visual scene. |
| Outcome: | The proposed method enables users to interact with the system using natural language instructions, allowing for a higher level of interactivity. |
Copied to clipboard
| Challenge: | Existing foundation models for general knowledge graph reasoning have focused on their structural aspects, with most efforts restricted to in-KG tasks. |
| Approach: | They propose a conditional encoding architecture that bridges the gap between textual and structural modalities, enabling seamless integration. |
| Outcome: | The proposed model outperforms baseline models on 28 datasets and is generalized to out-of-KG tasks. |
Copied to clipboard
| Challenge: | Building socially-intelligent AI agents involves creating agents that can sense, perceive, reason about, learn from, and respond to affect, behavior, and cognition of other agents. |
| Approach: | They propose a set of technical challenges and open questions for researchers to advance Social-AI. |
| Outcome: | The proposed frameworks are based on the social intelligence competencies that evolved over thousands of years in Homo sapiens and are expected to be the foundations for the development of social-intelligent AI agents. |
Copied to clipboard
| Challenge: | Current slice discovery methods in computer vision rely on converting input images into sets of attributes and testing hypotheses about configurations of pre-computed attributes associated with elevated error patterns. |
| Approach: | They propose a method to identify systematic biases in the mistakes of pre-trained vision models by converting input images into sets of attributes and testing hypotheses about configurations of these attributes. |
| Outcome: | The proposed method outperforms existing methods on 3 natural and 3 medical imaging datasets and generates pseudo-labels for each identified bias. |
Copied to clipboard
| Challenge: | BrainLoc is a lightweight object detection model guided by fMRI signals. |
| Approach: | They propose a brain-based object detection model guided by fMRI signals . they employ a multi-modal alignment strategy that enhances fmr feature extraction . |
| Outcome: | The proposed model improves fMRI-based object detection accuracy and convenience. |
Copied to clipboard
| Challenge: | AGL1K is the first audio geo-localization benchmark for audio language models (ALMs) it is based on a crowd-sourced platform and is available in 72 countries and territories. |
| Approach: | They propose a benchmark for audio geo-localization that quantifies the informativeness of each recording and a metric that quantizes the information of each audio clip. |
| Outcome: | The proposed benchmarks cover 72 countries and territories and can be used to improve audio geo-localization. |
Copied to clipboard
| Challenge: | a new benchmark for computer vision fails to capture richness and unpredictability of real-world anomalies . state-of-the-art VLMs struggle with visual anomaly perception and commonsense reasoning . elucidating the nature of anomalies is a fundamental human trait . |
| Approach: | They propose a benchmark for visual anomalies that includes annotations for visual grounding and categorizing anomalies based on their visual manifestations, their complexity, severity, and commonness. |
| Outcome: | The proposed benchmark improves on existing vision models by incorporating visual annotations. |